Papers with hypothesis testing problem
Bayes Test of Precision, Recall, and F1 Measure for Comparison of Two Natural Language Processing Models (P19-1)
Copied to clipboard
| Challenge: | Existing t-tests for cross-validation (CV) are inappropriate for model comparison . existing t tests for cross validation (CV), such as 52 CV t test and F ttest, are inadequate . |
| Approach: | They propose to use a block-regularized 32 CV to compare two NLP models . they calibrate the posterior distributions of P, R, and F1 and derive an accurate interval estimation of P and R . |
| Outcome: | The proposed model could regularize the difference in certain frequency distributions over linguistic units and yield stable estimators of P, R, and F1. |
Principled Detection of Hallucinations in Large Language Models via Multiple Testing (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to detect hallucinations are prone to generating false alarms and false feedbacks. |
| Approach: | They propose a method that aggregates multiple evaluation scores via conformal p-values, enabling calibrated detection with controlled false alarm rate. |
| Outcome: | The proposed method aggregates multiple evaluation scores via conformal p-values, enabling calibrated detection with controlled false alarm rate. |